Papers with open-source toolkit
PhoNLP: A joint multi-task learning model for Vietnamese part-of-speech tagging, named entity recognition and dependency parsing (2021.naacl-demos)
Copied to clipboard
| Challenge: | PhoNLP is a multi-task learning model for joint Vietnamese part-of-speech (POS) tagging, named entity recognition (NER) and dependency parsing. |
| Approach: | They propose a multi-task learning model for Vietnamese part-of-speech tagging, named entity recognition and dependency parsing that fine-tunes the pre-trained Vietnamese language model PhoBERT for each task independently. |
| Outcome: | The proposed model outperforms a single-task learning approach that fine-tunes the pre-trained Vietnamese language model PhoBERT for each task independently. |
Reliable, Reproducible, and Really Fast Leaderboards with Evalica (2025.coling-demos)
Copied to clipboard
| Challenge: | Using open-source evaluation tools, we create reliable and reproducible model leaderboards with human and machine feedback. |
| Approach: | They propose an open-source evaluation toolkit that facilitates the creation of reliable and reproducible model leaderboards. |
| Outcome: | The evaluation tool facilitates the creation of reliable and reproducible model leaderboards. |
NeurST: Neural Speech Translation Toolkit (2021.acl-demo)
Copied to clipboard
| Challenge: | a toolkit for speech translation is available for free and provides step-by-step recipes for feature extraction, data preprocessing, distributed training, and evaluation. |
| Approach: | They propose to use NeurST to facilitate speech translation research for NLP researchers . they show experimental results for different benchmark datasets which can be regarded as reliable baselines . |
| Outcome: | The proposed framework provides reliable benchmarks for speech translation research. |
MarkLLM: An Open-Source Toolkit for LLM Watermarking (2024.emnlp-demo)
Copied to clipboard
Leyi Pan, Aiwei Liu, Zhiwei He, Zitian Gao, Xuandong Zhao, Yijian Lu, Binglin Zhou, Shuliang Liu, Xuming Hu, Lijie Wen, Irwin King, Philip Yu
| Challenge: | Large Language Models (LLMs) embed imperceptible yet algorithmically detectable signals in outputs to identify LLM-generated text. |
| Approach: | They propose to develop an open-source toolkit for LLM watermarking that embeds imperceptible yet algorithmically detectable signals in model outputs to identify LLM-generated text. |
| Outcome: | MarkLLM provides a unified framework for implementing LLM watermarking algorithms, while providing user-friendly interfaces to ensure ease of access. |
OpenSLU: A Unified, Modularized, and Extensible Toolkit for Spoken Language Understanding (2023.acl-demo)
Copied to clipboard
| Challenge: | Spoken Language Understanding (SLU) is a task-oriented dialogue system . open-source toolkit provides a unified, modularized, and extensible toolkit for SLU . |
| Approach: | They introduce an open-source toolkit to provide a unified toolkit for spoken language understanding. |
| Outcome: | The proposed toolkit unifies 10 models for both single-intent and multi-intention scenarios. |
LeafNATS: An Open-Source Toolkit and Live Demo System for Neural Abstractive Text Summarization (N19-4)
Copied to clipboard
| Challenge: | Neural abstractive text summarization (NATS) has gained a lot of attention in the past few years from both industry and academia. |
| Approach: | They propose an open-source toolkit for training and evaluation of different sequence-to-sequence based models for the NATS task and for deploying the pre-trained models to real-world applications. |
| Outcome: | The proposed model can be used to generate high-quality summaries that are verbally innovative and can easily incorporate external knowledge. |
fastHan: A BERT-based Multi-Task Toolkit for Chinese NLP (2021.acl-demo)
Copied to clipboard
| Challenge: | Recently, the need for Chinese natural language processing (NLP) has a dramatic increase for many downstream applications. |
| Approach: | They propose to use Chinese word segmentation (CWS), Part-of-Speech (POS) tagging, named entity recognition (NER), and dependency parsing to train a multi-task model based on a pruned BERT. |
| Outcome: | The proposed model performs better than popular segmentation tools on a non-training corpus. |
LocalRQA: From Generating Data to Locally Training, Testing, and Deploying Retrieval-Augmented QA Systems (2024.acl-demos)
Copied to clipboard
| Challenge: | Existing tools for augmented question-answering do not support researchers and developers to customize the training, testing, and deployment process. |
| Approach: | They propose an open-source toolkit that features a wide selection of model training algorithms, evaluation methods, and deployment tools curated from the latest research. |
| Outcome: | The proposed framework trains and deploys 7B-models with the same performance as OpenAI’s text-ada-002 and GPT-4-turbo. |
RiskLab: A Controlled Toolkit for Probing Emergent Risks in LLM-Based Multi-Agent Systems (2026.acl-demo)
Copied to clipboard
Yu Jiang, Wenjie Wang, Yue Huang, Yanbo Wang, Zhenhong Zhou, Xiuying Chen, Yang Liu, Pin-Yu Chen, Wei Wang, Xiangliang Zhang
| Challenge: | Recent advances in large language model (LLM) agents have accelerated deployment of multi-agent systems for complex tasks. |
| Approach: | They propose an open-source toolkit for instantiating, probing, and measuring emergent risks in LLM-based multi-agent systems under controlled conditions. |
| Outcome: | The proposed toolkit is based on a structured topology–environment–protocol–agent–task quintuple enabling reproducible studies of how communication structure, coordination mechanisms, and incentives shape system-level risks. |
TokLens: A Multilingual Lens on Tokenizer Quality for LLMs (2026.acl-srw)
Copied to clipboard
| Challenge: | TokLens is an open-source toolkit for evaluating tokenizer quality across languages . authors evaluated 24 tokenizers from major LLM families across 15 typologically diverse languages - a gap that is stark in Japanese . |
| Approach: | They evaluate 24 tokenizers from major LLM families across 15 typologically diverse languages and correlate these metrics with downstream performance. |
| Outcome: | The proposed tokenizers produce 56x more tokens per word in Japanese than in English . the newer tokenizer Qwen2.5 and Gemma-2 reduce this gap to under 4x . |
AutoAlign: Get Your LLM Aligned with Minimal Annotations (2025.acl-demo)
Copied to clipboard
Xinyu Lu, Dong Xu, Chunkang Zhang, Xinyan Guan, Junxiang Wang, Qingyu Zhang, Pengbo Wang, Yingzhi Mao, Hao Xiang, Xueru Wen, Zichao Li, Yaojie Lu, Hongyu Lin, Le Sun, Xianpei Han
| Challenge: | Automated Alignment (ALM) is a set of algorithms designed to align Large Language Models (LLMs) with human intentions and values while minimizing manual intervention. |
| Approach: | They propose an open-source toolkit that integrates mainstream automated algorithms through a consistent interface and an accessible workflow supporting one-click execution for prompt synthesis and automatic alignment signal construction. |
| Outcome: | The proposed framework enables easy reproduction of existing results through extensive benchmarks and facilitates the development of novel approaches via modular components. |
ConvLab-2: An Open-Source Toolkit for Building, Evaluating, and Diagnosing Dialogue Systems (2020.acl-demos)
Copied to clipboard
Qi Zhu, Zheng Zhang, Yan Fang, Xiang Li, Ryuichi Takanobu, Jinchao Li, Baolin Peng, Jianfeng Gao, Xiaoyan Zhu, Minlie Huang
| Challenge: | ConvLab-2 inherits Convlab's framework but integrates more powerful dialogue models and supports more datasets. |
| Approach: | They present ConvLab-2, an open-source toolkit that enables researchers to build task-oriented dialogue systems with state-of-the-art models and perform an end-to-end evaluation. |
| Outcome: | The new tool inherits ConvLab's framework and extends it by integrating many recently proposed state-of-the-art dialogue models. |
NeuSpell: A Neural Spelling Correction Toolkit (2020.emnlp-demos)
Copied to clipboard
| Challenge: | a new spelling correction toolkit is available for free. |
| Approach: | They propose an open-source toolkit for spelling correction in English . they train neural models using spelling errors in context and using richer contextual representations. |
| Outcome: | The proposed spell-checker improves accuracy on synthetic examples and richer representations of the context. |
NeuroX Library for Neuron Analysis of Deep NLP Models (2023.acl-demo)
Copied to clipboard
| Challenge: | NeuroX is an open-source toolkit to conduct neuron analysis of natural language processing models. |
| Approach: | They propose a Python toolkit to conduct neuron analysis of natural language processing models. |
| Outcome: | a new open-source toolkit enables neuron analysis of natural language processing models . the framework provides a framework for data processing and evaluation, making it easier for researchers and practitioners to perform neuron analyses. |
CRSLab: An Open-Source Toolkit for Building Conversational Recommender System (2021.acl-demo)
Copied to clipboard
Kun Zhou, Xiaolei Wang, Yuanhang Zhou, Chenzhan Shang, Yuan Cheng, Wayne Xin Zhao, Yaliang Li, Ji-Rong Wen
| Challenge: | Existing studies on conversational recommender systems lack a unified and standardized implementation or comparison. |
| Approach: | They propose to use a unified framework and highly-decoupled modules to develop CRSs. |
| Outcome: | The proposed framework collects 6 commonly used human-annotated CRS datasets and implements 19 models that include advanced techniques such as graph neural networks and pre-training models. |
BMInf: An Efficient Toolkit for Big Model Inference and Tuning (2022.acl-demo)
Copied to clipboard
Xu Han, Guoyang Zeng, Weilin Zhao, Zhiyuan Liu, Zhengyan Zhang, Jie Zhou, Jun Zhang, Jia Chao, Maosong Sun
| Challenge: | Recent years, pre-trained language models (PLMs) have achieved promising results on various NLP tasks. |
| Approach: | They propose an open-source toolkit for big model inference and tuning which can support big model tuning at extremely low computation cost. |
| Outcome: | The proposed toolkit can support big model inference and tuning at extremely low computation cost. |
RobustQA: A Framework for Adversarial Text Generation Analysis on Question Answering Systems (2023.emnlp-demo)
Copied to clipboard
Yasaman Boreshban, Seyed Morteza Mirbostani, Seyedeh Fatemeh Ahmadi, Gita Shojaee, Fatemeh Kamani, Gholamreza Ghassem-Sani, Seyed Abolghasem Mirroshandel
| Challenge: | Question answering (QA) systems have reached human-level accuracy, but they are not robust enough and vulnerable to adversarial examples. |
| Approach: | They modified the attack algorithms widely used in text classification to fit them for QA systems. |
| Outcome: | The proposed framework is the first open-source toolkit for investigating textual adversarial attacks in QA systems. |
BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation (2026.acl-demo)
Copied to clipboard
| Challenge: | Existing document translation pipelines face a tension between linguistic processing and layout preservation. |
| Approach: | They propose a framework for layout-preserving PDF translation that decouples visual layout metadata from semantic content. |
| Outcome: | The proposed framework improves layout fidelity, visual aesthetics, and terminology consistency over representative baselines while maintaining competitive translation precision. |
Texar: A Modularized, Versatile, and Extensible Toolkit for Text Generation (P19-3)
Copied to clipboard
Zhiting Hu, Haoran Shi, Bowen Tan, Wentao Wang, Zichao Yang, Tiancheng Zhao, Junxian He, Lianhui Qin, Di Wang, Xuezhe Ma, Zhengzhong Liu, Xiaodan Liang, Wanrong Zhu, Devendra Sachan, Eric Xing
| Challenge: | Texar is an open-source text generation toolkit that supports a broad set of text generation tasks. |
| Approach: | They introduce Texar, an open-source text generation toolkit that supports text generation tasks. |
| Outcome: | Texar supports machine translation, summarization, dialog, content manipulation, and more. |
OpenT2T: An Open-Source Toolkit for Table-to-Text Generation (2024.emnlp-demo)
Copied to clipboard
Haowei Zhang, Shengyun Si, Yilun Zhao, Lujing Xie, Zhijian Xu, Lyuhao Chen, Linyong Nan, Pengcheng Wang, Xiangru Tang, Arman Cohan
| Challenge: | Existing methods for table-to-text generation are limited and benchmarked on a limited number of datasets. |
| Approach: | They propose to use open-source tools to reproduce existing large language models for performance comparison and expedite the development of new models. |
| Outcome: | The proposed toolkit compares existing large language models on 9 table-to-text generation datasets and maintains a leaderboard to provide insights for future work. |
Open-Theatre: An Open-Source Toolkit for LLM-based Interactive Drama (2025.emnlp-demos)
Copied to clipboard
| Challenge: | Existing tools for creating, modifying, and experimenting with interactive dramas are limited. |
| Approach: | They propose an open-source toolkit for creating configurable LLM-based interactive drama. |
| Outcome: | The proposed toolkit enhances narrative coherence and realistic behavior in interactions with agents. |
PyOpenDial: A Python-based Domain-Independent Toolkit for Developing Spoken Dialogue Systems with Probabilistic Rules (D19-3)
Copied to clipboard
| Challenge: | a recent development of spoken dialogue systems has enabled deep learning to achieve state-of-the-art performance. |
| Approach: | They propose a Python-based domain-independent, open-source toolkit for spoken dialogue systems. |
| Outcome: | The proposed toolkit extends OpenDial's Java-based architecture and provides new functions for neural dialogue state tracking and action planning. |
RAGViz: Diagnose and Visualize Retrieval-Augmented Generation (2024.emnlp-demo)
Copied to clipboard
| Challenge: | Large language models (LLMs) lack domain-specific knowledge and can cause hallucinations. |
| Approach: | They propose a RAG diagnosis tool that visualizes the attentiveness of the generated tokens in retrieved documents. |
| Outcome: | RAGViz provides token and document-level attention visualization and generation comparison upon context document addition and removal. |
LEGOEval: An Open-Source Toolkit for Dialogue System Evaluation via Crowdsourcing (2021.acl-demo)
Copied to clipboard
| Challenge: | Currently, researchers use automatic metrics and human evaluation to evaluate dialogue systems. |
| Approach: | They propose to use a Python API to easily evaluate dialogue systems using Amazon Mechanical Turk. |
| Outcome: | The open-source toolkit provides a fast, consistent method for reproducing human evaluation results. |
NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails (2023.emnlp-demo)
Copied to clipboard
| Challenge: | NeMo Guardrails is an open-source toolkit for easily adding programmable guardrails to LLM-based conversational systems. |
| Approach: | They propose to add programmable guardrails to LLMs that are user-defined, independent of the underlying LLM, and interpretable. |
| Outcome: | The proposed approach can be used with several LLM providers to develop controllable and safe LLM applications using programmable rails. |
KMatrix-2: A Comprehensive Heterogeneous Knowledge Collaborative Enhancement Toolkit for Large Language Model (2025.emnlp-demos)
Copied to clipboard
Shun Wu, Di Wu, Wangtao Sun, Ziyang Huang, Xiaowei Yuan, Kun Luo, XueYou Zhang, Shizhu He, Jun Zhao, Kang Liu
| Challenge: | Existing studies on K-LLMs systems focus on declarative knowledge and procedural knowledge (rules) . |
| Approach: | They propose to build a toolkit that supports comprehensive heterogeneous knowledge collaborative enhancement for Large Language Models (LLMs). |
| Outcome: | The proposed toolkit provides unified knowledge integration and joint knowledge retrieval methods to achieve more comprehensive heterogeneous knowledge collaborative enhancement. |
OpenICL: An Open-Source Framework for In-context Learning (2023.acl-demo)
Copied to clipboard
| Challenge: | In-context Learning (ICL) is a new paradigm for large language model evaluation. |
| Approach: | They propose an open-source toolkit for ICL and LLM evaluation. |
| Outcome: | The proposed framework is highly flexible and flexible and can be easily combined with other tools to suit users' needs. |
WIKIR: A Python Toolkit for Building a Large-scale Wikipedia-based English Information Retrieval Dataset (2020.lrec-1)
Copied to clipboard
| Challenge: | ad-hoc information retrieval methods usually require large amounts of annotated data to be effective. |
| Approach: | They propose an open-source toolkit to automatically build large-scale English information retrieval datasets based on Wikipedia. |
| Outcome: | The proposed toolkit builds large-scale English information retrieval datasets based on Wikipedia with 59,252 queries and 2,617,003 pairs. |
CSPB: Conversational Speech Processing Benchmark for Self-supervised Speech Models (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing benchmarks focus on clean, single-speaker, single channel audio, failing to reflect the complexities of natural human interaction. |
| Approach: | They propose a benchmark to assess the robustness of self-supervised speech models in conversational settings. |
| Outcome: | The proposed benchmark assesses the robustness of self-supervised speech models in conversational scenarios. |
Neural Network Models for Paraphrase Identification, Semantic Textual Similarity, Natural Language Inference, and Question Answering (C18-1)
Copied to clipboard
| Challenge: | Sentence pair modeling is a fundamental technique underlying many NLP tasks. |
| Approach: | They analyze several neural network designs for sentence pair modeling and compare their performance extensively across eight datasets. |
| Outcome: | The proposed models perform well across eight datasets including paraphrase identification, semantic textual similarity, natural language inference, and question answering tasks. |
Teaching Old Tokenizers New Words: Efficient Tokenizer Adaptation for Pretrained Models (2026.findings-eacl)
Copied to clipboard
| Challenge: | Extending existing vocabulary is a widely used step in adapting pre-trained language models to new domains or languages. |
| Approach: | They propose to extend a pre-trained tokenizer by continuing the BPE merge learning process on new data. |
| Outcome: | The proposed method improves tokenization efficiency and improves model utilization. |
LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers (2025.findings-naacl)
Copied to clipboard
Anton Razzhigaev, Matvey Mikhalchuk, Temurbek Rahmatullaev, Elizaveta Goncharova, Polina Druzhinina, Ivan Oseledets, Andrey Kuznetsov
| Challenge: | Large Language Models (LLMs) encode and store contextual information, but internal mechanisms are opaque. |
| Approach: | They propose a toolkit that assesses token-level nonlinearity, evaluates contextual memory, visualizes intermediate layer contributions and measures intrinsic dimensionality of representations. |
| Outcome: | The proposed framework assesses token-level nonlinearity, evaluates contextual memory, visualizes intermediate layer contributions, and measures the intrinsic dimensionality of representations. |
On the Evaluation of Speech Foundation Models for Spoken Language Understanding (2024.findings-acl)
Copied to clipboard
Siddhant Arora, Ankita Pasad, Chung-Ming Chien, Jionghao Han, Roshan Sharma, Jee-weon Jung, Hira Dhamyal, William Chen, Suwon Shon, Hung-yi Lee, Karen Livescu, Shinji Watanabe
| Challenge: | Spoken language understanding evaluation (SLUE) benchmarks are used to benchmark complex spoken language understanding tasks on natural speech. |
| Approach: | They propose a set of benchmark tasks to evaluate spoken language understanding on natural speech . they use pre-trained speech foundation models to evaluate the utility of different SFMs . |
| Outcome: | The proposed framework outperforms pre-trained speech foundation models on natural speech . the proposed framework also outperformed self-supervised SFMs on the sequence generation tasks . |
Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework (2026.findings-acl)
Copied to clipboard
Jiaqi Weng, Han Zheng, Hanyu Zhang, Ej Zhou, Qinqin He, Jialing Tao, Hui Xue, Zhixuan Chu, Xiting Wang
| Challenge: | Existing studies on how SAEs derive most fine-grained latent features for safety remain unexplored. |
| Approach: | They propose a framework for interpreting SAE features in safety-critical domains . they train a suite of SAEs with human-readable explanations and systematic evaluations based on pornography, politics, violence, and terror . |
| Outcome: | The proposed framework reduces interpretation cost by 55% and improves safety-critical features. |
skLEP: A Slovak General Language Understanding Benchmark (2025.findings-acl)
Copied to clipboard
Marek Suppa, Andrej Ridzik, Daniel Hládek, Tomáš Javůrek, Viktória Ondrejová, Kristína Sásiková, Martin Tamajka, Marian Simko
| Challenge: | skLEP is the first comprehensive benchmark specifically designed for evaluating Slovak natural language understanding models. |
| Approach: | They introduce a benchmark specifically designed for evaluating Slovak natural language understanding models. |
| Outcome: | The proposed benchmark covers nine tasks that span token-level, sentence-pair, document-level tasks. |